CSDV3017  ·  DEVOPS  ·  SCHOOL OF COMPUTER SCIENCE, UPES

The Resilience Process:
Detect, Alert, Respond, Refine

Lecture 9 — turning the resiliency patterns from last class into a repeatable operational loop that catches failure, reacts to it, and gets a little smarter every time.

InstructorDr. Mohsin Furkh Dar
SessionWeek 3 · Tue, 23 Jun 2026
Time14:00 – 15:00
UnitUnit III
I
II
III
IV
V
VI
VII
WHERE WE LEFT OFF

Recap & today's agenda

Lecture 8 · Done

Architecture & Resiliency

  • Monolithic vs microservices architecture
  • How architecture shapes DevOps pipelines & ownership
  • Resiliency patterns: redundancy, auto-scaling, circuit breakers...
  • Why cloud infrastructure makes resiliency affordable
Lecture 9 · Today

The Resilience Process

  • The four-stage loop: Detect → Alert → Respond → Refine
  • What happens, and which tools live, at each stage
  • Key metrics: MTTD, MTTA, MTTR, MTBF
  • Common pitfalls — and how the loop ties back to Kaizen
FROM PATTERNS TO PROCESS

Patterns alone don't save you at 3 a.m.

"Hope is not a strategy — but neither is a circuit breaker nobody is watching." — why resiliency needs a process, not just patterns

Redundancy and failover only help if a human — or an automated system — notices the failure, reacts to it, and learns from it. That's exactly what the resilience process gives us.

THE BIG PICTURE

A continuous loop, not a checklist

CONTINUOUS RESILIENCE LOOP Detect Alert Respond Refine
EACH INCIDENT FEEDS THE NEXT CYCLE

Four stages, one purpose

The resilience process is the operational habit that makes redundancy, failover and auto-scaling actually work in practice — a loop a team runs every time something goes wrong, so the system gets steadily harder to break.

  • Detect — notice something is wrong
  • Alert — get the right people informed, fast
  • Respond — act to restore service
  • Refine — learn, and close the gap for next time
STAGE 1 OF 4

Detect — notice it before your users do

1 · Detect2 · Alert3 · Respond4 · Refine

What's happening here

Continuous monitoring of metrics, logs, and traces watches for signs that something has drifted from normal — latency spikes, error rates, failed health checks.

What good detection needs

  • Clear baselines — you must know "normal" to spot "wrong"
  • Metrics, logs, and distributed traces working together
  • Synthetic checks that test the user's actual journey
STAGE 2 OF 4

Alert — get it to the right human, fast

1 · Detect2 · Alert3 · Respond4 · Refine

What's happening here

A detected anomaly is turned into an alert: it's routed, prioritised, and pushed to whoever is on-call — through chat, paging apps, or phone calls for the most severe cases.

What good alerting needs

  • Sensible thresholds — too sensitive causes alert fatigue
  • Clear severity levels, so urgency is obvious at a glance
  • A defined on-call rotation with backup escalation
STAGE 3 OF 4

Respond — restore service, then breathe

1 · Detect2 · Alert3 · Respond4 · Refine

What's happening here

The on-call engineer (or an automated system) acts to restore service — a rollback, a failover, a restart, or a scripted remediation — guided by a runbook where one exists.

What good response needs

  • Runbooks for known failure modes, written before the fire
  • Automated remediation for the most common, well-understood issues
  • Calm, blameless communication while the incident is live
STAGE 4 OF 4

Refine — close the loop

1 · Detect2 · Alert3 · Respond4 · Refine

What's happening here

Once service is restored, the team runs a blameless postmortem: what failed, why detection or response was slow, and what concrete change prevents a repeat.

This is Kaizen, applied to incidents

Recall Lecture 7: Kaizen is continuous, incremental improvement made by the people doing the work. Refine is that habit applied directly to reliability — every incident makes the system a little harder to break next time.

THE TOOLKIT

Tools, mapped to each stage

Stage Category Example tools
Detect Metrics, logs & tracing Prometheus, Grafana, ELK Stack, Datadog
Alert Paging & notification PagerDuty, Opsgenie, Slack/Teams integrations
Respond Runbooks & automation Ansible, scripted rollback, ChatOps bots
Refine Postmortem & tracking Shared incident docs, Jira/GitLab Tracker action items

Tool names matter less than the discipline: each stage needs an owner, even if the "tool" is a shared document.

MEASURING THE LOOP

Metrics that matter

MTTD

Mean Time to Detect

How long from failure starting to it being noticed.

MTTA

Mean Time to Acknowledge

How long from alert firing to a human acknowledging it.

MTTR

Mean Time to Recover

How long from failure to service being fully restored.

MTBF

Mean Time Between Failures

How long, on average, the system runs before the next failure.

A healthy resilience process drives MTTD, MTTA and MTTR down over time — and pushes MTBF up. That trend is the real evidence that "Refine" is working.

WHERE THE LOOP BREAKS

Common pitfalls

Detect & Alert
  • Alert fatigue — too many low-value alerts, so real ones get ignored
  • No clear ownership — an alert fires, but nobody is sure whose job it is
  • Monitoring the server, not the user's actual experience
Respond & Refine
  • No runbooks — every incident relearns the same lessons under pressure
  • Blame-driven postmortems that make people hide problems
  • Action items from postmortems that are written down and never done
CLASS DISCUSSION

Walk the loop on a real scenario

Scenario

Checkout service starts timing out at peak traffic

A microservice from Lecture 8's "Orders" service starts timing out under heavy load. Walk through all four stages: what would Detect catch first, who should Alert reach, what's a reasonable first Response, and what's one concrete Refine action item for next sprint?

D
What signal fires first?
A
Who gets paged?
R
First safe action?
R
One fix for next time?
UNIT III, SO FAR

How today connects back

Lecture 7

Adoption & Kaizen

Refine is Kaizen applied directly to incidents — small, continuous fixes after every failure.

Lecture 8

Architecture & Resiliency

Redundancy, circuit breakers and failover are the raw material the Respond stage actually uses.

Lecture 9

The Resilience Process

The operational loop that makes all of the above real, every single day, on call.

WRAP-UP

Today, in three lines

  • Resilience is a loop — Detect, Alert, Respond, Refine — not a one-time setup.
  • Each stage needs clear ownership, the right tooling, and a defined process — not just good intentions.
  • MTTD, MTTA, MTTR and MTBF tell you if the loop is actually improving over time.
Next lecture · Lecture 10

Mon, 29 Jun 2026 · 14:00–15:00 · Unit IV

DevOps principles; Version Control (SVN, Git, GitHub); Gitflow workflow; CI with GitHub Actions.

Before next class

Quick prep

Think of a time something broke and nobody noticed right away — which stage of this loop was missing?

CSDV3017 · DEVOPS
SHEET 01/14